incidents: add INC-132, Deadbugz MCP call-count-gated tool poisoning - #109
Conversation
|
Thanks for this — the record is careful work, and the call-count gating detail is a genuinely useful thing to have in the index. Two things, one of them my fault. 1. The next free id is Happy to do the renumber and rebase for you if you'd rather not — say the word and I'll push it to your branch, or open a follow-up that does it, whichever you prefer. Only asking first because it's your branch. I'm also adding a guard so this can't recur: a validator check that fails on duplicate incident ids, plus a 2. Triage verdict: it qualifies. I checked it against One thing worth a second look while you're rebasing: the record sets Nothing else blocks this from my side. |
Pillar Security's Deadbugz disclosure: a malicious MCP server serves its documented, benign tool contract for the first two tool calls in a session, then substitutes credential-harvesting instructions on the third call. A connect-and-check review never crosses the threshold, so this is invisible to single-session auditing by construction. Renumbered from INC-132 to INC-136: GenAI-Security-Project#117 claimed 132-134 while this was open, GenAI-Security-Project#122 (open) claims 135. Category is research-demonstrated, not real-world: the schema defines real-world as a confirmed incident, and this entry's own impact field says no compromise was confirmed, 23 delivery PRs opened, none merged. INC-126 is the one other entry in the dataset with comparable unconfirmed-impact language and it carries the same category. Mapped to ASI01 (goal hijack), ASI02 (tool misuse), ASI04 (agentic supply chain). Source: https://www.pillar.security/blog/deadbugz-currently-active-mcp-supply-chain-campaign
2a06d79 to
d9d3394
Compare
|
Thanks for catching both. Pushed a fix for each. Renumbered to INC-136 and rebuilt clean off current main. #117 and #122 both touch the same generated JSON/JS this schema regenerates, and a hand-resolved merge there risked corrupting it. Checked the free id against main directly: 001 through 134 contiguous, 135 claimed by your open #122, so 136 stays clear either way that one lands. On category, I went with your suggestion, research-demonstrated. Wanted to write down why it's right, since I checked it rather than took it on faith. The schema's own enum description defines real-world as a confirmed incident, and this entry's impact field says plainly that no compromise was confirmed, 23 delivery PRs, none merged. INC-126 is the only other entry in the dataset with comparable unconfirmed-impact language, and it already carries research-demonstrated. Three independent readings landing on the same answer is what convinced me. One small thing on your review: I couldn't find docs/TRIAGE_RULES.md anywhere in the repo, on main or any open PR. Not blocking anything here, the schema enum backs the same verdict on its own. Flagging in case it's sitting uncommitted somewhere on your end. |
|
Thanks — both changes check out, and the category reasoning from the schema enum is the right basis. You're right about Verification of this branch (
One non-blocking suggestion. The convention for
Swapping it in is your call. Either way, the failure stays a draft ( Merge-order note, so you aren't surprised later. This PR and #122 (INC-135) each add an incident and regenerate the same files ( CI hasn't run here. Workflows on fork PRs wait for a maintainer to approve them, so the green results above are local runs, not GitHub checks. |
…te one (#124) INC-132 was allocated twice - by #117 and by #109 - because both read the end of data/incidents.json while the other was open, and the second one only found out when the merge conflicted. A duplicate id also silently breaks anything that resolves an incident by id: the webapp deep link, the evidence join, the reports. validate.js gains checkIncidentIds(), which fails on a duplicate and names it. scripts/next-incident-id.mjs prints a free id; --check-prs also accounts for ids claimed by open pull requests, which is what would have caught this one - it currently reports INC-136, because INC-135 is claimed by open PR #122. The duplicate case is deliberately not tested by mutating data/incidents.json: node --test runs suites in parallel, so writing to the shared corpus races the suites reading it, which is a bug this repository has already had. The guard is covered by a corpus-uniqueness test and a wiring test instead, and was negative-tested by hand: injecting a duplicate made validate.js exit 1 with "INC-006 is used by 2 records". Baseline before: 0 errors, 88 warnings, 327 passed; 85/85 tests. Baseline after : 0 errors, 88 warnings, 328 passed; 89/89 tests, twice. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…-behaviour incidents (#122) * Fail on duplicate incident ids, and give contributors a way to allocate one INC-132 was allocated twice - by #117 and by #109 - because both read the end of data/incidents.json while the other was open, and the second one only found out when the merge conflicted. A duplicate id also silently breaks anything that resolves an incident by id: the webapp deep link, the evidence join, the reports. validate.js gains checkIncidentIds(), which fails on a duplicate and names it. scripts/next-incident-id.mjs prints a free id; --check-prs also accounts for ids claimed by open pull requests, which is what would have caught this one - it currently reports INC-136, because INC-135 is claimed by open PR #122. The duplicate case is deliberately not tested by mutating data/incidents.json: node --test runs suites in parallel, so writing to the shared corpus races the suites reading it, which is a bug this repository has already had. The guard is covered by a corpus-uniqueness test and a wiring test instead, and was negative-tested by hand: injecting a duplicate made validate.js exit 1 with "INC-006 is used by 2 records". Baseline before: 0 errors, 88 warnings, 327 passed; 85/85 tests. Baseline after : 0 errors, 88 warnings, 328 passed; 89/89 tests, twice. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Add INC-135 (llmware SQLi), and record what kind of failure a CVE record is CVE-2026-85689: llmware 0.4.6 interpolates filter and lookup values straight into SQL WHERE clauses in both backends, so an attacker-controlled filter neutralises the caller's scoping and returns rows from other documents and collections; on PostgreSQL it is full boolean/UNION injection. The record states its status plainly rather than implying a tidy disclosure: the CNA is VulnCheck rather than the vendor, the 7.1 is VulnCheck's own secondary metric, llmware has published no advisory, the upstream report (llmware-ai/llmware#1304, opened 2026-06-12) is still open with no maintainer response, and no release after the affected 0.4.6 exists. It is recorded as unfixed. **No control_failures.** The instruction was to add one only where a vendor advisory explicitly names a control that failed or was absent. llmware has published no advisory, so no such statement exists and the array is omitted rather than filled from the reporter's words. Schema change — two new optional fields on Incident: incident_class tooling-cve | ai-behaviour. A conventional software vulnerability in GenAI tooling is AI-relevant because of what it exposes, not because the model misbehaved; the index should not blur the two. INC-132..134 are backfilled as tooling-cve, which is what they are. mapping_status draft | sme-confirmed. Says who decided owasp_entries, not anything about the incident. INC-132..135 are all draft: an agent proposed the entries and no SME has signed them. Records predating the fields simply lack them; nothing else is backfilled. Baseline before and after: 0 errors, 88 warnings, 327 passed; 85/85 tests. Incidents 134 -> 135. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
#122 (INC-135) and #109 (INC-136) each regenerated the incident counts against a main that lacked the other, so after both landed data/stats.json, the README stats and docs/incidents.js were one short: stats:check failed and two unit tests failed (136 !== 135). Regenerated with generate.js + npm run stats. Removed the "pillar-security" tag from INC-136: no other incident tags the disclosing vendor, and the source is already cited in references and external_refs. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…136) (#179) #122 (INC-135) and #109 (INC-136) each added an incident and regenerated the count from their own base, so main ended at 136 incidents with every generated total reading 135. Content Validation and two unit tests failed on 5d67ee7 ("136 !== 135"). Output of generate.js, npm run stats and render-stats.mjs; no hand edits. Three lines: README stats marker, data/stats.json, docs/incidents.js header. Co-authored-by: sim <sim@local> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Pillar Security's Deadbugz disclosure: a malicious MCP server serves
its documented, benign tool contract for the first two tool calls in
a session, then substitutes credential-harvesting instructions on the
third call. A connect-and-check review never crosses the threshold,
so this is invisible to single-session auditing by construction.
Mapped to ASI01 (goal hijack), ASI02 (tool misuse), ASI04 (agentic
supply chain). Delivery: 23 PRs from one account against unrelated
repos, none merged at time of disclosure.
Source: https://www.pillar.security/blog/deadbugz-currently-active-mcp-supply-chain-campaign